Papers with GPT models
Reconstruct Your Previous Conversations! Comprehensively Investigating Privacy Leakage Risks in Conversations with GPT Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing GPT models allow users to interact with them for multiple rounds to optimize the task execution. |
| Approach: | They propose a conversation reconstruction attack targeting the contents of previous conversations between GPT models and benign users, i.e., the benign users’ input contents during their interaction with GPT. |
| Outcome: | The proposed attacks demonstrate that GPT-4's defense mechanisms are ineffective against these attacks. |
CAPE: Context-Aware Personality Evaluation Framework for Large Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies use a context-free approach to assess humans . existing studies use the Disney World test, which ignores real-world applications . |
| Approach: | They propose a framework to assess personality traits in large language models . they use conversational history to quantify the consistency of LLM responses . |
| Outcome: | The proposed framework improves consistency of responses in large language models . it also shows that conversational history enhances consistency and personality shifts . |
Experimental versus In-Corpus Variation in Referring Expression Choice (2024.lrec-main)
Copied to clipboard
| Challenge: | Against our expectations, the divergence is greatest between the corpus and the GPT model. |
| Approach: | They compare the results of three studies to examine how well the corpus can model variation . they find that experimental methodology introduces substantial noise . |
| Outcome: | The results show that the corpus can model variation captured from the corpuse and RE form choices made during experiments. |
Can LLMs replace Neil deGrasse Tyson? Evaluating the Reliability of LLMs as Science Communicators (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) and AI assistants are experiencing exponential growth in usage among expert and amateur users. |
| Approach: | They propose to assess the reliability of current Large Language Models as science communicators . they use a dataset comprising 742 Yes/No queries embedded in complex scientific concepts . |
| Outcome: | The proposed model outperforms open-access models in scientific question-answering tasks . the model outpersforms GPT-4 Turbo models in many evaluation aspects . |